Papers by Jelmer van der Linde

2 papers
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .
The EuroPat Corpus: A Parallel Corpus of European Patent Data (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of patent-specific parallel data is available for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Approach: They present a patent-specific corpus of parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Outcome: The filtered corpus ranges in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpus being up to 2 times larger.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations